NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

Fair active learning

https://doi.org/10.1016/j.eswa.2022.116981

Anahideh, Hadis; Asudeh, Abolfazl; Thirumuruganathan, Saravanan (August 2022, Expert Systems with Applications)

Full Text Available
Shahin: Faster Algorithms for Generating Explanations for Multiple Predictions

https://doi.org/10.1145/3448016.3457332

Hasani, Sona; Thirumuruganathan, Saravanan; Koudas, Nick; Das, Gautam (June 2021, ACM SIGMOD)
null (Ed.)
Full Text Available
Astrid: accurate selectivity estimation for string predicates using deep learning

https://doi.org/10.14778/3436905.3436907

Shetiya, Suraj; Thirumuruganathan, Saravanan; Koudas, Nick; Das, Gautam (December 2020, Proceedings of the VLDB Endowment)
null (Ed.)
Accurate selectivity estimation for string predicates is a long-standing research challenge in databases. Supporting pattern matching on strings (such as prefix, substring, and suffix) makes this problem much more challenging, thereby necessitating a dedicated study. Traditional approaches often build pruned summary data structures such as tries followed by selectivity estimation using statistical correlations. However, this produces insufficiently accurate cardinality estimates resulting in the selection of sub-optimal plans by the query optimizer. Recently proposed deep learning based approaches leverage techniques from natural language processing such as embeddings to encode the strings and use it to train a model. While this is an improvement over traditional approaches, there is a large scope for improvement. We propose Astrid, a framework for string selectivity estimation that synthesizes ideas from traditional and deep learning based approaches. We make two complementary contributions. First, we propose an embedding algorithm that is query-type (prefix, substring, and suffix) and selectivity aware. Consider three strings 'ab', 'abc' and 'abd' whose prefix frequencies are 1000, 800 and 100 respectively. Our approach would ensure that the embedding for 'ab' is closer to 'abc' than 'abd'. Second, we describe how neural language models could be used for selectivity estimation. While they work well for prefix queries, their performance for substring queries is sub-optimal. We modify the objective function of the neural language model so that it could be used for estimating selectivities of pattern matching queries. We also propose a novel and efficient algorithm for optimizing the new objective function. We conduct extensive experiments over benchmark datasets and show that our proposed approaches achieve state-of-the-art results.
more » « less
Full Text Available
On Benchmarking for Crowdsourcing and Future of Work Platform

Borromeo, Ria Mae; Chen, Lei; Dubey, Abhishek; Roy, Sudeepa; Thirumuruganathan, Saravanan (January 2020, A Quarterly bulletin of the Computer Society of the IEEE Technical Committee on Data Engineering)

Online crowdsourcing platforms have proliferated over the last few years and cover a number of important domains, these platforms include worker-task platforms such as Amazon Mechanical Turk, worker-for hire platforms such as TaskRabbit to specialized platforms with specific tasks such as ridesharing like Uber, Lyft, Ola, etc. An increasing proportion of the human workforce will be employed by these platforms in the near future. The crowdsourcing community has done yeoman’s work in designing effective algorithms for various key components, such as incentive design, task assignment, and quality control. Given the increasing importance of these crowdsourcing platforms, it is now time to design mechanisms so that it is easier to evaluate the effectiveness of these platforms. Specifically, we advocate developing benchmarks for crowdsourcing research. Benchmarks often identify important issues for the community to focus on and improve upon. This has played a key role in the development of research domains as diverse as databases and deep learning. We believe that developing appropriate benchmarks for crowdsourcing will ignite further innovations. However, crowdsourcing – and future of work, in general – is a very diverse field that makes developing benchmarks much more challenging. Substantial effort is needed that spans across developing benchmarks for datasets, metrics, algorithms, platforms, and so on. In this article, we initiate some discussion into this important problem and issue a call-to-arms for the community to work on this important initiative.
more » « less
Full Text Available
Efficient Signal Reconstruction for a Broad Range of Applications

https://doi.org/10.1145/3371316.3371327

Asudeh, Abolfazl; Augustine, Jees; Nazi, Azade; Thirumuruganathan, Saravanan; Zhang, Nan; Das, Gautam; Srivastava, Divesh (November 2019, ACM SIGMOD Record)
null (Ed.)
Full Text Available
Orca-SR: A Real-Time Traffic Engineering Framework leveraging Similarity Joins

Augustine, Jees; Shetiya, Suraj; Asudeh, Azadeh; Thirumuruganathan, Saravanan; Nazi, Azade; Zhang, Nan; Das, Gautam; Srivastava, Divesh (January 2020, Proceedings of the VLDB Endowment)
null (Ed.)
Full Text Available
EXPLAINER: Entity Resolution Explanations

https://doi.org/10.1109/ICDE.2019.00224

Ebaid, Amr; Thirumuruganathan, Saravanan; Aref, Walid G.; Elmagarmid, Ahmed; Ouzzani, Mourad (April 2019, 2019 IEEE 35th International Conference on Data Engineering (ICDE)x)

Full Text Available
Optimized group formation for solving collaborative tasks

https://doi.org/10.1007/s00778-018-0516-7

Rahman, Habibur; Roy, Senjuti Basu; Thirumuruganathan, Saravanan; Amer-Yahia, Sihem; Das, Gautam (February 2019, The VLDB Journal)

Full Text Available
A Human-in-the-loop Attribute Design Framework for Classification

https://doi.org/10.1145/3308558.3313547

Salam, Md Abdus; Koone, Mary E.; Thirumuruganathan, Saravanan; Das, Gautam; Basu Roy, Senjuti (January 2019, The World Wide Web Conference, {WWW} 2019)

Full Text Available
Efficient construction of approximate ad-hoc ML models through materialization and reuse

https://doi.org/10.14778/3236187.3236199

Hasani, Sona; Thirumuruganathan, Saravanan; Asudeh, Abolfazl; Koudas, Nick; Das, Gautam (July 2018, Proceedings of the VLDB Endowment)

Full Text Available

« Prev Next »

Search for: All records